lowram: Keep w1 in packed form during signature generation - #1345
lowram: Keep w1 in packed form during signature generation#1345mkannwischer wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Mac Mini (M1, 2020) benchmarks (opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
46502 cycles |
46534 cycles |
1.00 |
ML-DSA-44 sign |
131401 cycles |
131164 cycles |
1.00 |
ML-DSA-44 verify |
47327 cycles |
47342 cycles |
1.00 |
ML-DSA-65 keypair |
81701 cycles |
81716 cycles |
1.00 |
ML-DSA-65 sign |
215390 cycles |
215441 cycles |
1.00 |
ML-DSA-65 verify |
79329 cycles |
79334 cycles |
1.00 |
ML-DSA-87 keypair |
132436 cycles |
132496 cycles |
1.00 |
ML-DSA-87 sign |
277552 cycles |
277435 cycles |
1.00 |
ML-DSA-87 verify |
133580 cycles |
133546 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Mac Mini (M1, 2020) benchmarks (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
112924 cycles |
112777 cycles |
1.00 |
ML-DSA-44 sign |
403184 cycles |
401308 cycles |
1.00 |
ML-DSA-44 verify |
119532 cycles |
119390 cycles |
1.00 |
ML-DSA-65 keypair |
193857 cycles |
192926 cycles |
1.00 |
ML-DSA-65 sign |
651728 cycles |
649924 cycles |
1.00 |
ML-DSA-65 verify |
193025 cycles |
193010 cycles |
1.00 |
ML-DSA-87 keypair |
318845 cycles |
318891 cycles |
1.00 |
ML-DSA-87 sign |
831120 cycles |
828752 cycles |
1.00 |
ML-DSA-87 verify |
321851 cycles |
321752 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
AMD EPYC 3rd gen (c6a)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
52243 cycles |
52005 cycles |
1.00 |
ML-DSA-44 sign |
156689 cycles |
155314 cycles |
1.01 |
ML-DSA-44 verify |
54300 cycles |
54379 cycles |
1.00 |
ML-DSA-65 keypair |
89232 cycles |
89891 cycles |
0.99 |
ML-DSA-65 sign |
252045 cycles |
255341 cycles |
0.99 |
ML-DSA-65 verify |
89122 cycles |
89591 cycles |
0.99 |
ML-DSA-87 keypair |
143951 cycles |
142178 cycles |
1.01 |
ML-DSA-87 sign |
310297 cycles |
310461 cycles |
1.00 |
ML-DSA-87 verify |
139225 cycles |
139090 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton5
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
51564 cycles |
51457 cycles |
1.00 |
ML-DSA-44 sign |
162252 cycles |
161886 cycles |
1.00 |
ML-DSA-44 verify |
54672 cycles |
54811 cycles |
1.00 |
ML-DSA-65 keypair |
90155 cycles |
90177 cycles |
1.00 |
ML-DSA-65 sign |
266116 cycles |
268015 cycles |
0.99 |
ML-DSA-65 verify |
89823 cycles |
89977 cycles |
1.00 |
ML-DSA-87 keypair |
145851 cycles |
145957 cycles |
1.00 |
ML-DSA-87 sign |
335809 cycles |
335685 cycles |
1.00 |
ML-DSA-87 verify |
145427 cycles |
144930 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
AMD EPYC 4th gen (c7a)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
47026 cycles |
46740 cycles |
1.01 |
ML-DSA-44 sign |
139819 cycles |
145382 cycles |
0.96 |
ML-DSA-44 verify |
49294 cycles |
51476 cycles |
0.96 |
ML-DSA-65 keypair |
82252 cycles |
82956 cycles |
0.99 |
ML-DSA-65 sign |
227024 cycles |
229646 cycles |
0.99 |
ML-DSA-65 verify |
82074 cycles |
82585 cycles |
0.99 |
ML-DSA-87 keypair |
130391 cycles |
129744 cycles |
1.00 |
ML-DSA-87 sign |
280889 cycles |
280564 cycles |
1.00 |
ML-DSA-87 verify |
129532 cycles |
128041 cycles |
1.01 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A76 (Raspberry Pi 5) benchmarks (opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
112309 cycles |
112045 cycles |
1.00 |
ML-DSA-44 sign |
354052 cycles |
353393 cycles |
1.00 |
ML-DSA-44 verify |
117053 cycles |
117282 cycles |
1.00 |
ML-DSA-65 keypair |
194750 cycles |
194574 cycles |
1.00 |
ML-DSA-65 sign |
583585 cycles |
583697 cycles |
1.00 |
ML-DSA-65 verify |
192910 cycles |
193189 cycles |
1.00 |
ML-DSA-87 keypair |
320601 cycles |
320560 cycles |
1.00 |
ML-DSA-87 sign |
747320 cycles |
747028 cycles |
1.00 |
ML-DSA-87 verify |
318394 cycles |
318140 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton5 (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
100288 cycles |
100193 cycles |
1.00 |
ML-DSA-44 sign |
369472 cycles |
369630 cycles |
1.00 |
ML-DSA-44 verify |
110068 cycles |
110334 cycles |
1.00 |
ML-DSA-65 keypair |
176009 cycles |
175189 cycles |
1.00 |
ML-DSA-65 sign |
596526 cycles |
597631 cycles |
1.00 |
ML-DSA-65 verify |
177530 cycles |
177767 cycles |
1.00 |
ML-DSA-87 keypair |
286614 cycles |
286470 cycles |
1.00 |
ML-DSA-87 sign |
758280 cycles |
759318 cycles |
1.00 |
ML-DSA-87 verify |
295698 cycles |
296310 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
AMD EPYC 3rd gen (c6a) (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
133444 cycles |
133190 cycles |
1.00 |
ML-DSA-44 sign |
520204 cycles |
518479 cycles |
1.00 |
ML-DSA-44 verify |
146679 cycles |
146665 cycles |
1.00 |
ML-DSA-65 keypair |
223925 cycles |
224540 cycles |
1.00 |
ML-DSA-65 sign |
847147 cycles |
844646 cycles |
1.00 |
ML-DSA-65 verify |
234409 cycles |
234394 cycles |
1.00 |
ML-DSA-87 keypair |
367457 cycles |
367714 cycles |
1.00 |
ML-DSA-87 sign |
1062489 cycles |
1061302 cycles |
1.00 |
ML-DSA-87 verify |
380562 cycles |
381497 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Intel Xeon 4th gen (c7i)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
43230 cycles |
43292 cycles |
1.00 |
ML-DSA-44 sign |
130048 cycles |
130069 cycles |
1.00 |
ML-DSA-44 verify |
45248 cycles |
45143 cycles |
1.00 |
ML-DSA-65 keypair |
75909 cycles |
75764 cycles |
1.00 |
ML-DSA-65 sign |
213758 cycles |
213624 cycles |
1.00 |
ML-DSA-65 verify |
74418 cycles |
74335 cycles |
1.00 |
ML-DSA-87 keypair |
123062 cycles |
122992 cycles |
1.00 |
ML-DSA-87 sign |
271631 cycles |
271188 cycles |
1.00 |
ML-DSA-87 verify |
120766 cycles |
120668 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Intel Xeon 3rd gen (c6i)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
61412 cycles |
61655 cycles |
1.00 |
ML-DSA-44 sign |
189046 cycles |
188845 cycles |
1.00 |
ML-DSA-44 verify |
66231 cycles |
66290 cycles |
1.00 |
ML-DSA-65 keypair |
109628 cycles |
112698 cycles |
0.97 |
ML-DSA-65 sign |
312172 cycles |
318478 cycles |
0.98 |
ML-DSA-65 verify |
109663 cycles |
111904 cycles |
0.98 |
ML-DSA-87 keypair |
170175 cycles |
171213 cycles |
0.99 |
ML-DSA-87 sign |
379730 cycles |
379008 cycles |
1.00 |
ML-DSA-87 verify |
170594 cycles |
170855 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
AMD EPYC 4th gen (c7a) (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
118667 cycles |
118359 cycles |
1.00 |
ML-DSA-44 sign |
457628 cycles |
458607 cycles |
1.00 |
ML-DSA-44 verify |
131501 cycles |
129909 cycles |
1.01 |
ML-DSA-65 keypair |
201384 cycles |
200886 cycles |
1.00 |
ML-DSA-65 sign |
740800 cycles |
744497 cycles |
1.00 |
ML-DSA-65 verify |
208939 cycles |
208873 cycles |
1.00 |
ML-DSA-87 keypair |
331948 cycles |
331103 cycles |
1.00 |
ML-DSA-87 sign |
936713 cycles |
938146 cycles |
1.00 |
ML-DSA-87 verify |
343584 cycles |
345239 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton4
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
67496 cycles |
67139 cycles |
1.01 |
ML-DSA-44 sign |
198509 cycles |
198230 cycles |
1.00 |
ML-DSA-44 verify |
70148 cycles |
70284 cycles |
1.00 |
ML-DSA-65 keypair |
119095 cycles |
119469 cycles |
1.00 |
ML-DSA-65 sign |
325504 cycles |
326490 cycles |
1.00 |
ML-DSA-65 verify |
116584 cycles |
116939 cycles |
1.00 |
ML-DSA-87 keypair |
196377 cycles |
196410 cycles |
1.00 |
ML-DSA-87 sign |
421551 cycles |
421431 cycles |
1.00 |
ML-DSA-87 verify |
193276 cycles |
192954 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Intel Xeon 4th gen (c7i) (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
92216 cycles |
91733 cycles |
1.01 |
ML-DSA-44 sign |
349576 cycles |
351109 cycles |
1.00 |
ML-DSA-44 verify |
99758 cycles |
99446 cycles |
1.00 |
ML-DSA-65 keypair |
154324 cycles |
153920 cycles |
1.00 |
ML-DSA-65 sign |
567577 cycles |
571127 cycles |
0.99 |
ML-DSA-65 verify |
159730 cycles |
160330 cycles |
1.00 |
ML-DSA-87 keypair |
255364 cycles |
255096 cycles |
1.00 |
ML-DSA-87 sign |
725239 cycles |
721477 cycles |
1.01 |
ML-DSA-87 verify |
263819 cycles |
264366 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Intel Xeon 3rd gen (c6i) (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
154136 cycles |
154277 cycles |
1.00 |
ML-DSA-44 sign |
588309 cycles |
588660 cycles |
1.00 |
ML-DSA-44 verify |
168888 cycles |
169423 cycles |
1.00 |
ML-DSA-65 keypair |
262317 cycles |
267849 cycles |
0.98 |
ML-DSA-65 sign |
958311 cycles |
979945 cycles |
0.98 |
ML-DSA-65 verify |
271809 cycles |
277179 cycles |
0.98 |
ML-DSA-87 keypair |
431619 cycles |
432934 cycles |
1.00 |
ML-DSA-87 sign |
1211290 cycles |
1215004 cycles |
1.00 |
ML-DSA-87 verify |
446915 cycles |
447425 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A76 (Raspberry Pi 5) benchmarks (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
211677 cycles |
211596 cycles |
1.00 |
ML-DSA-44 sign |
761225 cycles |
759367 cycles |
1.00 |
ML-DSA-44 verify |
228934 cycles |
229045 cycles |
1.00 |
ML-DSA-65 keypair |
375171 cycles |
375197 cycles |
1.00 |
ML-DSA-65 sign |
1245358 cycles |
1247878 cycles |
1.00 |
ML-DSA-65 verify |
371103 cycles |
371441 cycles |
1.00 |
ML-DSA-87 keypair |
600704 cycles |
600145 cycles |
1.00 |
ML-DSA-87 sign |
1585805 cycles |
1584430 cycles |
1.00 |
ML-DSA-87 verify |
616453 cycles |
615831 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton4 (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
128011 cycles |
127916 cycles |
1.00 |
ML-DSA-44 sign |
441684 cycles |
441335 cycles |
1.00 |
ML-DSA-44 verify |
136368 cycles |
136365 cycles |
1.00 |
ML-DSA-65 keypair |
223088 cycles |
221710 cycles |
1.01 |
ML-DSA-65 sign |
714012 cycles |
714085 cycles |
1.00 |
ML-DSA-65 verify |
220581 cycles |
220606 cycles |
1.00 |
ML-DSA-87 keypair |
364536 cycles |
365306 cycles |
1.00 |
ML-DSA-87 sign |
914776 cycles |
916248 cycles |
1.00 |
ML-DSA-87 verify |
371022 cycles |
370929 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton3
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
71530 cycles |
71457 cycles |
1.00 |
ML-DSA-44 sign |
209907 cycles |
208907 cycles |
1.00 |
ML-DSA-44 verify |
74817 cycles |
75005 cycles |
1.00 |
ML-DSA-65 keypair |
125984 cycles |
125992 cycles |
1.00 |
ML-DSA-65 sign |
344118 cycles |
345143 cycles |
1.00 |
ML-DSA-65 verify |
124046 cycles |
124143 cycles |
1.00 |
ML-DSA-87 keypair |
206428 cycles |
206682 cycles |
1.00 |
ML-DSA-87 sign |
439635 cycles |
444029 cycles |
0.99 |
ML-DSA-87 verify |
204505 cycles |
204086 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton3 (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
138537 cycles |
138455 cycles |
1.00 |
ML-DSA-44 sign |
486610 cycles |
486114 cycles |
1.00 |
ML-DSA-44 verify |
149235 cycles |
149248 cycles |
1.00 |
ML-DSA-65 keypair |
241159 cycles |
242184 cycles |
1.00 |
ML-DSA-65 sign |
790421 cycles |
791620 cycles |
1.00 |
ML-DSA-65 verify |
241516 cycles |
241529 cycles |
1.00 |
ML-DSA-87 keypair |
395244 cycles |
396160 cycles |
1.00 |
ML-DSA-87 sign |
1011999 cycles |
1013640 cycles |
1.00 |
ML-DSA-87 verify |
404107 cycles |
403824 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton2
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
112444 cycles |
112659 cycles |
1.00 |
ML-DSA-44 sign |
354470 cycles |
354364 cycles |
1.00 |
ML-DSA-44 verify |
117205 cycles |
117531 cycles |
1.00 |
ML-DSA-65 keypair |
195775 cycles |
195098 cycles |
1.00 |
ML-DSA-65 sign |
587160 cycles |
584870 cycles |
1.00 |
ML-DSA-65 verify |
194231 cycles |
193581 cycles |
1.00 |
ML-DSA-87 keypair |
320708 cycles |
321321 cycles |
1.00 |
ML-DSA-87 sign |
747632 cycles |
747338 cycles |
1.00 |
ML-DSA-87 verify |
318503 cycles |
318952 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Graviton2 (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
212383 cycles |
212400 cycles |
1.00 |
ML-DSA-44 sign |
761080 cycles |
760457 cycles |
1.00 |
ML-DSA-44 verify |
229961 cycles |
229630 cycles |
1.00 |
ML-DSA-65 keypair |
376776 cycles |
375851 cycles |
1.00 |
ML-DSA-65 sign |
1247361 cycles |
1248638 cycles |
1.00 |
ML-DSA-65 verify |
372562 cycles |
371900 cycles |
1.00 |
ML-DSA-87 keypair |
602074 cycles |
602038 cycles |
1.00 |
ML-DSA-87 sign |
1586617 cycles |
1591124 cycles |
1.00 |
ML-DSA-87 verify |
618303 cycles |
618524 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A72 (Raspberry Pi 4) benchmarks (opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
218089 cycles |
213397 cycles |
1.02 |
ML-DSA-44 sign |
599197 cycles |
587414 cycles |
1.02 |
ML-DSA-44 verify |
217607 cycles |
213073 cycles |
1.02 |
ML-DSA-65 keypair |
379744 cycles |
376687 cycles |
1.01 |
ML-DSA-65 sign |
986692 cycles |
979172 cycles |
1.01 |
ML-DSA-65 verify |
364473 cycles |
363486 cycles |
1.00 |
ML-DSA-87 keypair |
639246 cycles |
647130 cycles |
0.99 |
ML-DSA-87 sign |
1311852 cycles |
1349859 cycles |
0.97 |
ML-DSA-87 verify |
619801 cycles |
632424 cycles |
0.98 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A72 (Raspberry Pi 4) benchmarks (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
300220 cycles |
297307 cycles |
1.01 |
ML-DSA-44 sign |
1133706 cycles |
1124831 cycles |
1.01 |
ML-DSA-44 verify |
331808 cycles |
327013 cycles |
1.01 |
ML-DSA-65 keypair |
549132 cycles |
547240 cycles |
1.00 |
ML-DSA-65 sign |
1869739 cycles |
1861858 cycles |
1.00 |
ML-DSA-65 verify |
530042 cycles |
527234 cycles |
1.01 |
ML-DSA-87 keypair |
849544 cycles |
841613 cycles |
1.01 |
ML-DSA-87 sign |
2351348 cycles |
2350913 cycles |
1.00 |
ML-DSA-87 verify |
880873 cycles |
873805 cycles |
1.01 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A55 (Snapdragon 888) benchmarks (opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
273189 cycles |
271602 cycles |
1.01 |
ML-DSA-44 sign |
813853 cycles |
811152 cycles |
1.00 |
ML-DSA-44 verify |
273807 cycles |
273547 cycles |
1.00 |
ML-DSA-65 keypair |
465037 cycles |
466624 cycles |
1.00 |
ML-DSA-65 sign |
1332507 cycles |
1341019 cycles |
0.99 |
ML-DSA-65 verify |
453524 cycles |
452548 cycles |
1.00 |
ML-DSA-87 keypair |
799137 cycles |
798726 cycles |
1.00 |
ML-DSA-87 sign |
1842572 cycles |
1835055 cycles |
1.00 |
ML-DSA-87 verify |
787003 cycles |
773446 cycles |
1.02 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
Arm Cortex-A55 (Snapdragon 888) benchmarks (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
466035 cycles |
465745 cycles |
1.00 |
ML-DSA-44 sign |
2138837 cycles |
2140830 cycles |
1.00 |
ML-DSA-44 verify |
556780 cycles |
558445 cycles |
1.00 |
ML-DSA-65 keypair |
783328 cycles |
784167 cycles |
1.00 |
ML-DSA-65 sign |
3498055 cycles |
3497460 cycles |
1.00 |
ML-DSA-65 verify |
868982 cycles |
867417 cycles |
1.00 |
ML-DSA-87 keypair |
1266853 cycles |
1272563 cycles |
1.00 |
ML-DSA-87 sign |
4330173 cycles |
4331369 cycles |
1.00 |
ML-DSA-87 verify |
1392739 cycles |
1392417 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
There was a problem hiding this comment.
SpacemiT K1 8 (Banana Pi F3) benchmarks (no-opt)
Details
| Benchmark suite | Current: 7c8238d | Previous: 41c00da | Ratio |
|---|---|---|---|
ML-DSA-44 keypair |
757718 cycles |
760195 cycles |
1.00 |
ML-DSA-44 sign |
3137744 cycles |
3141418 cycles |
1.00 |
ML-DSA-44 verify |
857097 cycles |
859027 cycles |
1.00 |
ML-DSA-65 keypair |
1288623 cycles |
1288559 cycles |
1.00 |
ML-DSA-65 sign |
5092434 cycles |
5082523 cycles |
1.00 |
ML-DSA-65 verify |
1368044 cycles |
1368131 cycles |
1.00 |
ML-DSA-87 keypair |
2112748 cycles |
2109622 cycles |
1.00 |
ML-DSA-87 sign |
6377925 cycles |
6359264 cycles |
1.00 |
ML-DSA-87 verify |
2226700 cycles |
2225228 cycles |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
CBMC Results (ML-DSA-87, REDUCE-RAM)
Full Results (212 proofs)
|
CBMC Results (ML-DSA-65, REDUCE-RAM)
Full Results (212 proofs)
|
CBMC Results (ML-DSA-87)
Full Results (212 proofs)
|
CBMC Results (ML-DSA-44)
Full Results (212 proofs)
|
There was a problem hiding this comment.
This is an interesting and impactful optimization, thank you for proposing it @mkannwischer.
Overall, I think this is worth implementing.
However, we are diverging more and more from the spec and reference implementation here. Especially with the lazy/eager split for yvec, we have to take utmost care to keep the documentation of both source and its relation to FIPS-204 accurate and crisp in terms of providing exactly the information that is needed to reason about the code at any point -- and nothing more to, to avoid unnecessary cognitive/reasoning load. I left a few comments in this direction.
We should cite the relevant papers. |
49c22a9 to
cffde5c
Compare
The signing attempt held w1 as a full mld_polyveck, but w1 is only
needed as the w1Encode input to H and, coefficient-wise, as the a1
argument of MakeHint. Both work from the packed encoding, so Decompose
and w1Encode now run one polynomial at a time into a
MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers
each row with the new mld_polyw1_unpack.
With w1 gone, the matrix-vector scratch is the sole remaining user of
the buffer the two shared. In REDUCE_RAM mode that scratch is a single
polynomial rather than a polyvecl, and it now shares storage with z, so
the buffer disappears.
Signing allocation in bytes:
default REDUCE_RAM
ML-DSA-44 44704 -> 44448 13120 -> 9792
ML-DSA-65 69312 -> 68032 17248 -> 11872
ML-DSA-87 108224 -> 107200 21344 -> 14176
mld_polyveck_decompose and mld_polyveck_pack_w1 have no other callers
and are replaced by the fused mld_polyveck_decompose_pack_w1.
Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
cffde5c to
c8b0bd4
Compare
| void mld_polyvec_matrix_pointwise_montgomery_yvec_eager( | ||
| mld_polyveck *w, mld_polymat_eager *mat, const mld_yvec_eager *y, | ||
| mld_yvec_scratch_eager *scratch) |
There was a problem hiding this comment.
Have you considered whether the scratch can be embedded into mld_yvec, which anyway is an abstract type? Needs some checking of lifetimes in sign.c.
This would simplify the interface.
There was a problem hiding this comment.
I don't see a clean way to do that. It would require a union inside of mld_yvec (or pay extra 1 KiB), and you would have to change mld_compute_pack_z. This may lead to issues with the CBMC proofs.
I'd prefer to keep it as is.
| requires(memory_no_alias(r, sizeof(mld_poly))) | ||
| requires(memory_no_alias(a, MLDSA_POLYW1_PACKEDBYTES_88)) | ||
| assigns(memory_slice(r, sizeof(mld_poly))) | ||
| ensures(array_bound(r->coeffs, 0, MLDSA_N, 0, 1 << 6)) |
There was a problem hiding this comment.
MLDSA_POLYW1_PACKED_BITS_88 ?
| requires(memory_no_alias(r, sizeof(mld_poly))) | ||
| requires(memory_no_alias(a, MLDSA_POLYW1_PACKEDBYTES_32)) | ||
| assigns(memory_slice(r, sizeof(mld_poly))) | ||
| ensures(array_bound(r->coeffs, 0, MLDSA_N, 0, 1 << 4)) |
There was a problem hiding this comment.
MLDSA_POLYW1_PACKED_BITS_32 ?
| void mld_polyvec_matrix_pointwise_montgomery_yvec_lazy(mld_polyveck *w, | ||
| mld_polymat_lazy *mat, |
There was a problem hiding this comment.
ISTM that the introduction of mld_yvec_scratch_{lazy/eager} is a preparatory refactor that can stand alone. Seeing the large complexity of the commit already, can you please hoist this into a separate commit?
| for (i = 0; i < MLDSA_N / 4; ++i) | ||
| __loop__( | ||
| invariant(i <= MLDSA_N/4) | ||
| invariant(array_bound(r->coeffs, 0, i*4, 0, 1 << 6)) |
There was a problem hiding this comment.
MLDSA_POLYW1_PACKED_BITS_88
| for (i = 0; i < MLDSA_N / 2; ++i) | ||
| __loop__( | ||
| invariant(i <= MLDSA_N/2) | ||
| invariant(array_bound(r->coeffs, 0, i*2, 0, 1 << 4)) |
There was a problem hiding this comment.
MLDSA_POLYW1_PACKED_BITS_32
| { | ||
| unsigned int i; | ||
|
|
||
| for (i = 0; i < MLDSA_N / 2; ++i) |
There was a problem hiding this comment.
Inconsistent use of ++i vs prevailing i++ in the rest of the code base
| { | ||
| unsigned int i; | ||
|
|
||
| for (i = 0; i < MLDSA_N / 4; ++i) |
There was a problem hiding this comment.
Inconsistent use of ++i vs prevailing i++ in the rest of the code base
| #define MLDSA_POLYW1_PACKED_BITS_88 6 | ||
| #define MLDSA_POLYW1_PACKED_BITS_32 4 |
There was a problem hiding this comment.
Introducing those is slightly inconsistent with existing packing routines, which hardcode their bit widths. E.g. mld_polyt1_[un]pack.
| #if !defined(MLD_CONFIG_NO_SIGN_API) | ||
| #if defined(MLD_CONFIG_MULTILEVEL_WITH_SHARED) || MLD_CONFIG_PARAMETER_SET == 44 | ||
| MLD_INTERNAL_API | ||
| void mld_polyw1_unpack_88(mld_poly *r, |
There was a problem hiding this comment.
I realize that those have no use prior to the final optimization, but can you still hoist the introduction of the w1 unpacking routines and their CBMC proofs into a separate commit?
hanno-becker
left a comment
There was a problem hiding this comment.
I am unhappy with the complexity creep, but have so far not found any approach to avoiding it.
In the meantime, some notes:
- Can you split the work into multiple commits? As indicated in the comments, you can at least hoist the yvec scratch refactor and the w1_unpack changes into standalone commits.
- The use of
MLDSA_POLYW1_PACKED_BITS_XXis inconsistent, and its presence inconsistent with other pack bit width, e.g. fort1.
The signing attempt held w1 as a full mld_polyveck, but w1 is only needed as the w1Encode input to H and, coefficient-wise, as the a1 argument of MakeHint. Both work from the packed encoding, so Decompose and w1Encode now run one polynomial at a time into a MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers each row with the new mld_polyw1_unpack.
With w1 gone, the matrix-vector scratch is the sole remaining user of the buffer the two shared. In REDUCE_RAM mode that scratch is a single polynomial rather than a polyvecl, and it now shares storage with z, so the buffer disappears.
Signing allocation in bytes:
mld_polyveck_decompose and mld_polyveck_pack_w1 have no other callers and are replaced by the fused mld_polyveck_decompose_pack_w1.